Papers with semantic encoding

5 papers
Exploring Conditional Variational Mechanism to Pinyin Input Method for Addressing One-to-Many Mappings in Low-Resource Scenarios (2024.acl-short)

Copied to clipboard

Challenge: Experimental results demonstrate the superior performance of our method.
Approach: They propose to leverage conditional variational mechanism to simplify pinyin IME . they employ a strategy that facilitates interaction between pinyan and Chinese character information .
Outcome: The proposed method improves the performance of pinyin input method engine (IME) under low-resource conditions.
Using active learning to expand training data for implicit discourse relation recognition (D18-1)

Copied to clipboard

Challenge: Existing methods to determine semantic relations between text spans are limited in the field of discourse-level relation recognition.
Approach: They propose to expand the training data set using the corpus of explicitly-related arguments by arbitrarily dropping the overtly presented discourse connectives.
Outcome: The proposed model expands the training data set using the corpus of explicitly-related arguments, by arbitrarily dropping the overtly presented discourse connectives.
Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models have enabled their use as semantic encoders for recommendation, but their roles and behaviors in this setting are still not well understood.
Approach: They propose a benchmark to evaluate large language models as semantic encoders in recommendation scenarios.
Outcome: The proposed benchmark shows that ranking of 11 leading LLMs is low compared to MTEB, highlighting the unique challenges of semantic encoding in recommendation.
TabEmb: Joint Semantic-Structure Embedding for Table Annotation (2026.acl-long)

Copied to clipboard

Challenge: Existing tables learn by linearizing the 2D table into a 1D token sequence and encoding it with pretrained language models (PLMs) such as BERT, but this leads to limited semantic quality and weaker generalization to unseen or rare values compared to modern LLMs.
Approach: They propose a table annotation module called TabEmb which decouples semantic encoding from structural modeling by creating semantically rich embeddings for each column.
Outcome: Experiments show that TabEmb outperforms baselines on different table annotation tasks.
The Emergence of Semantic Units in Massively Multilingual Models (2024.lrec-main)

Copied to clipboard

Challenge: Massively multilingual models can process text in several languages relying on a shared set of parameters, but little is known about the encoding of multilingual information in single network units.
Approach: They propose to use a shared set of parameters to encode multilingual information in single network units.
Outcome: The proposed model achieves higher scores in semantic encoding in languages with more cross-lingual alignment than those with more shared cross-linguistic substrate.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations